Skip to content

feat(dashboard): add read-only fleet kanban - #1657

Closed
HelloWorldSungin wants to merge 70 commits into
kunchenguid:mainfrom
HelloWorldSungin:fm/dashboard-kanban-mvp
Closed

HelloWorldSungin wants to merge 70 commits into
kunchenguid:mainfrom
HelloWorldSungin:fm/dashboard-kanban-mvp

Conversation

@HelloWorldSungin

Copy link
Copy Markdown

Intent

Build story 1/6 of the Firstmate dashboard epic: a read-only, loopback-only mobile-first kanban SPA and small server that consume the landed fm-fleet-snapshot.v1 contract from issue #18 and the populated forge-agnostic work-item references from #21. The dashboard must never independently parse data/, state/, or projects/, must use each task card column and action exactly as computed by the contract precedence, and must contain no forge logic or fleet lifecycle controls. Run at most one fixed argument-based fm-fleet-snapshot.sh --json child at a time, debounce filesystem events and coalesce poll triggers, apply a hard timeout, safely bound stderr, retain last-known-good data with explicit freshness and error metadata, and guarantee a fleet change appears within one poll interval without reload. Serve only numeric loopback addresses, push updates with SSE and backoff reconnect, and render task id/title, project, kind, harness, model, effort, state detail, full PR URL, endpoint liveness, last-event age, and work-item links; show a plain link when enrichment is unavailable or the forge unsupported, and render no affordance at all when no reference exists. Keep persistent secondmates in a separate lane and provide accessible phone-width filters for project, harness, model, kind, and state. Gracefully represent empty, first-run, malformed-version, timeout, command-missing, and stale-last-good states. Package a boot-persistent user-level systemd service configurable by FM_HOME, numeric loopback address, port, poll interval, timeout, and stale threshold, with no sudo. Match the recorded Board prototype layout and hierarchy across wide and narrow phone treatments while choosing a maintainable implementation stack. Filesystem monitoring must prove zero writes under Firstmate data/, state/, and projects/, and stopping the dashboard must have zero effect on supervision. Delivery must ultimately be a PR on HelloWorldSungin/firstmate with Closes #11, a substantive issue #11 summary comment, and, if the known pipeline defect opens upstream, the upstream PR closed and the fork PR hand-raised with the literal no-mistakes marker plus an honest defect note.

What Changed

  • Add a mobile-first, read-only fleet kanban that renders contract-provided columns and actions, persistent secondmates, task metadata, work-item links, endpoint health, and accessible filters.
  • Add a loopback-only server with serialized, bounded snapshot refreshes, last-known-good error handling, SSE updates, reconnect backoff, and timely stale-state transitions.
  • Add a configurable no-sudo user-level systemd installer and document dashboard setup, configuration, and runtime behavior.

Risk Assessment

✅ Low: Captain, the authorized fixes close all three prior failure paths without introducing a material source-level regression, and the dashboard remains read-only and well bounded.

Testing

Inspected the targeted change, ran the focused dashboard behavior test, manually exercised the real browser surface and filtering at desktop and phone widths, captured rendered evidence and a live API envelope, verified dashboard shutdown leaves supervision untouched, and cleaned temporary fixtures. All checks passed.

  • Evidence: Populated wide kanban with secondmate lane and contract-derived cards (local file: /tmp/no-mistakes-evidence/01KZ5WM9ZRCRNYGSDQ8JQF7G5R/dashboard-fixture-wide-cards.png)
  • Evidence: Phone-width card with full PR URL, action, liveness, age, and enriched work item (local file: /tmp/no-mistakes-evidence/01KZ5WM9ZRCRNYGSDQ8JQF7G5R/dashboard-fixture-phone-full-card.png)
  • Evidence: Accessible phone-width filter layout (local file: /tmp/no-mistakes-evidence/01KZ5WM9ZRCRNYGSDQ8JQF7G5R/dashboard-fixture-phone.png)
Evidence: Live dashboard API contract envelope
{"phase":"ready","stale":false,"poll_seconds":1,"tasks":[{"id":"dashboard-task","column":"active","action":"supervise","work_item_count":1},{"id":"quiet-task","column":"waiting","action":"recheck","work_item_count":0},{"id":"deckhand","column":"secondmate","action":"route_work","work_item_count":0}]}
Evidence: Dashboard shutdown isolation check
dashboard stopped; independent supervision sentinel remained alive

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 3 issues found → auto-fixed ✅
  • 🚨 assets/dashboard/app.js:213 - Intent requires the dashboard to “use each task card column and action exactly as computed by the contract precedence,” but taskCard consumes only card.column; card.action is never read anywhere in the SPA. Render the literal contract-provided action as read-only card metadata, or explicitly approve omitting it.
  • 🚨 assets/dashboard/app.js:188 - Intent requires endpoint liveness, but endpointTone returns green whenever exists is true before checking for status === "dead". The snapshot contract explicitly permits a dead secondmate agent whose endpoint still exists, so such a secondmate gets a green card/lane indicator; the health summary also counts it as live. Make status authoritative and use endpoint existence only when liveness is unknown.
  • 🚨 bin/fm-dashboard-server.mjs:291 - The configurable stale threshold is not reflected when time alone makes a snapshot stale. After the initial fetch, the SPA relies on SSE, but heartbeats carry only comments and no broadcast occurs when staleMs is crossed. For example, poll=60s and stale=10s leaves the board showing “fresh” for about 50 seconds past the threshold. Schedule a freshness broadcast at the threshold or update freshness from heartbeat/client time.

🔧 Fix: Captain, correct dashboard liveness, actions, and stale transitions
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Inspected git diff --stat 8967ff1eccb241bb7a799c019fe76bb7d35fc192..7586569f7361ef1822c1887169b522ac2e86db08 and the dashboard implementation and tests.
  • bash tests/fm-dashboard.test.sh
  • Started the controlled dashboard fixture and rendered it in Chrome at 1440x1000 and 390x844.
  • Applied the harness filter through the rendered form and confirmed one matching card, zero secondmates, and (1 active).
  • Checked chrome-devtools-axi console for browser errors.
  • Queried /api/snapshot with curl and jq to confirm ready freshness plus contract-provided columns and actions.
  • Stopped the dashboard while an independent supervision sentinel was running and verified the sentinel remained alive.
  • Confirmed git status --short remained clean and transient fixture directories were removed.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Sungin Kim and others added 30 commits July 23, 2026 03:59
Add cursor-agent (Composer) and agy (Antigravity/Gemini) as CREW-ONLY,
HERDR-ONLY crewmate harnesses (captain-approved divergence; design in
data/captain.md, empirical verification in data/cursor-agy-verify/report.md).
They are never a primary runtime and never a secondmate launcher, and fm-spawn
fail-closes both gates before any backend or worktree work.

Launch construction moves into bin/fm-launch-lib.sh (extracted verbatim from
fm-spawn, protected by the existing fm-spawn-dispatch-profile regression) so the
per-harness template and model/effort renderers are unit-testable:
- cursor: `cursor-agent --trust --force <model> "$(cat brief)"`; effort is
  encoded in the parameterized model string (composer-2.5[effort=high]), so
  there is no standalone effort flag.
- agy: `agy --dangerously-skip-permissions <model> --effort <low|medium|high>
  --prompt-interactive "$(cat brief)"`.

agy workspace trust: an interactive agy launch gates on a per-workspace trust
modal that --dangerously-skip-permissions does not cover, and trust is an
exact-path entry in agy's shared global settings. bin/fm-agy-trust-lib.sh seeds
the exact worktree path before launch (fm-spawn, fail-closed) and removes it at
teardown (fm-teardown, main and secondmate-child paths); it is locked, atomic,
idempotent, and leaves an unparseable settings file untouched.

On herdr, agent-state, liveness, turn-end, and send-safety are all generic
(native `agent get` + the pane.agent_status_changed event stream), so no adapter
code and no turn-end hook are added, and no repo .cursor/hooks.json /
.agents/hooks.json is ever written. Composer stays `unknown` (the safe default);
no generic bare-glyph rule is added (agy's `>` glyph would be a dead-shell send
hazard). tmux is documented out-of-scope.

cursor/agy are added to the verified-adapter set in fm-bootstrap and
fm-dispatch-select (agy effort low|medium|high; cursor no standalone effort);
both are intentionally unscorable in quota-balanced selection (paid subs not
tracked by quota-axi) and fall through to the uniform random fallback.

Docs: harness-adapters SKILL and AGENTS.md section 4.

Tests: colocated unit tests for the launch lib and trust lib; contract tests for
the crew-only/herdr-only refusals and teardown trust cleanup; extended
dispatch-select/bootstrap validation; and a skip-guarded live herdr-lab smoke
that launches both CLIs through the real templates and confirms native
detection. Full suite green; shellcheck 0.11.0 clean.
…e-path tests

Address the independent adversarial review (data/cursor-agy-review/report.md):

B1 - raw launch-command bypass: the raw escape hatch derived the harness from the
first non-assignment word's basename (`cursor-agent`/`env`), so the cursor/agy
crew-only/herdr-only guard (keyed on the exact tokens `cursor`/`agy`) missed it.
Add fm_launch_raw_restricted_harness (resolves the real executable through `env`
and leading VAR=val assignments, with an all-words backstop) and refuse any raw
command that resolves to cursor-agent/agy outright - the raw path cannot provide
the trust seed or native supervision they require, so callers use --harness.

B2 - turn-end wake: the shared transition policy deliberately DEFERS herdr
idle/done, and cursor/agy install no hook and write no status file, so a
completed turn only ever surfaced via the unreliable content-hash stale
heuristic. Add a native-identity-gated, DEBOUNCED completion detector:
fm_transition_native_completion (pure, in fm-transition-lib.sh) plus
fm-watch.sh's maybe_native_turnend, which reads the pane's native agent_status
and touches state/<id>.turn-ended once a cursor/agy pane settles at idle/done
across two consecutive polls - the same completion signal every other harness's
hook writes, handled by the existing scan/wake path with no change to the shared
policy or event stream. Adds fm_backend_agent_status (raw enum accessor).

B3 - agy trust is now ownership-aware and transactional: fm_agy_trust_add reports
created vs preexisting so firstmate removes ONLY an entry it created and never
deletes a path the captain already trusted; the durable marker state/<id>.agy-trust
drives ownership-aware teardown; the spawn abort trap rolls back a created entry;
and teardown removal failure is an INCOMPLETE teardown that retains the marker +
metadata for a deterministic retry (same contract on the secondmate-child path).

B4 - the age-only 10s stale-lock breaker could steal a LIVE lock; replace it with
the repository's ownership-and-liveness lock (fm_lock_*), which never reclaims a
lock whose holder PID is alive and only lets the owner release.

S1 - add the missing failure-path tests: raw-command bypass variants (direct,
env, assignment-prefixed) on forbidden dimensions; created-vs-preexisting
ownership; parallel mutation; a live-lock-held-past-10s theft test; spawn-abort
rollback and teardown-incomplete-on-failure; and the live smoke now observes a
working->idle/done transition and PROVES a native completion wake for both CLIs.

Docs updated (harness-adapters skill, fm-launch-lib/fm-spawn comments) to describe
the debounced native completion detector instead of the earlier inaccurate
"native event stream drives turn-end" claim. Full suite green; shellcheck 0.11.0
clean. Pre-existing flake noted: tests/fm-watcher-lock.test.sh (unrelated beacon
timing; reproduces identically on the pre-change fm-watch.sh).
Match the repo convention (nearly all tests/*.test.sh are mode 0755) so direct
./tests/<name>.test.sh invocation works, not only via the bash runner.
Address the adversarial re-review (data/cursor-agy-rereview/report.md):

B1 (still-open shell-wrapper bypass) - a quoted wrapper like `bash -lc 'agy ...'`
or `sh -c "cursor-agent ..."` split into tokens `'agy` / `"cursor-agent`, so the
raw-guard's basename scan missed it and fm-spawn recorded harness=bash on tmux,
bypassing the crew-only/herdr-only gates. fm_launch_raw_restricted_harness now
adds a fail-closed backstop that neutralizes shell quote characters and command
separators before scanning every token's basename, so a wrapped cursor-agent/agy
invocation is caught and refused. Over-approximates safely (the raw hatch is for
unverified adapters; cursor/agy have a canonical --harness path).

B3 (still-open rollback leaks) - two fixes:
  1. AGY_TRUST_ROLLBACK_PATH is now armed BEFORE the fallible ownership-marker
     write, so a marker-write failure after the global trust entry was created
     still leaves the abort trap armed to roll it back.
  2. The rollback is factored into fm_agy_trust_rollback: on removal success it
     drops the marker; on removal FAILURE it (re)writes the marker with the
     leaked path as retry evidence - even when the marker was never written -
     instead of deleting it. The abort trap calls this tested function.

S2 - fm-bootstrap now emits `CREW_DISPATCH: backend mismatch - ...` when a valid
dispatch profile selects the crew-only herdr-only cursor/agy harness while the
resolved backend is not herdr, instead of deferring the surprise to spawn-time
refusal. bootstrap-diagnostics skill + AGENTS.md section 13 recognize the new
line.

Tests (would fail if unfixed): shell-wrapper bypass at both the resolver unit
level and through fm-spawn refusal; rollback clean-removal, rollback-removal
failure preserving marker evidence, and rollback failure recreating an absent
marker (the marker-write-failure evidence path); bootstrap backend-mismatch
warning on tmux and silence on herdr.

Full suite green; bin/fm-lint.sh (ShellCheck 0.11.0) exit 0. Not pushed.
Note: the end-to-end fm-spawn abort path reaches the agy seed only through a real
herdr spawn (no clean deterministic failure-injection seam), so the abort-trap
logic is covered by the fm_agy_trust_rollback unit tests that exercise the exact
function the trap runs.
…tests

B1 (variable-indirection raw shell wrappers, second re-review) - a raw command is
executed by the crewmate PANE's shell, so `AGY=agy bash -lc '$AGY ...'`,
`bash -lc 'x=agy; eval "$x ..."'`, backticks, `$(...)`, and split-token
concatenation (`$a$g`) can all expand to a restricted executable that a basename
scan cannot see. fm_launch_raw_restricted_harness now returns `unresolved` for any
raw command containing shell expansion ($ or backtick), and fm-spawn refuses it:
the command cannot be statically proven not to launch cursor/agy. Over-approximates
safely (the raw hatch is for unverified adapters; cursor/agy have a canonical
--harness path). Literal and env/assignment forms still resolve to cursor/agy.

Regressed-test fixes (full-suite failures caused by earlier commits in this
branch):
- tests/fm-gotmp.test.sh: teardown now sources bin/fm-agy-trust-lib.sh (which
  pulls in bin/fm-wake-lib.sh for its ownership lock); the fake teardown bin did
  not symlink them, so the source aborted teardown. Symlink both as newly required
  siblings, and make fm-agy-trust-lib's on-demand fm-wake-lib source tolerate an
  absent file (the lock is only exercised for agy tasks, whose real environment
  always has it).
- tests/fm-pi-watch-extension.test.sh: the pi secondmate launch TEMPLATE moved
  into bin/fm-launch-lib.sh during the launch-template extraction, so the test's
  grep of bin/fm-spawn.sh missed it; read both files (the placeholder substitution
  and tracked-extension path stay in fm-spawn.sh).

Tests: variable-indirection bypass rows at the resolver unit level and through the
fm-spawn refusal path. bin/fm-lint.sh (ShellCheck 0.11.0) clean. Not pushed.
These are pre-existing, environment-specific failures surfaced by running the full
suite in an en_US.UTF-8 host with Pi 0.81.1 and node in /usr/bin. None is caused by
the cursor/agy work; each is fixed at its true root cause, not masked.

- bin/fm-test-run.sh: the coverage guard sorts its lane/family lists with
  `LC_ALL=C sort` but ran `comm` in the ambient locale, so in en_US.UTF-8 comm
  rejected the C-collated input ("not in sorted order") and returned non-zero.
  Run every comm under LC_ALL=C to match its sort. (The cursor/agy test filenames
  happened to expose this latent locale inconsistency.)

- tests/fm-calm-pi-extension.test.sh: hard-pinned Pi 0.80.10; the host has 0.81.1.
  Empirically re-verified 2026-07-23 - the renderer and interactive-E2E assertions
  pass green against @earendil-works/pi-coding-agent 0.81.1 - so the pin now accepts
  0.80.10 or 0.81.1 (evidence-based, still fails loudly on an unverified version).

- tests/fm-session-start.test.sh: forced a MISSING diagnostic by removing `node`
  from its fake bin, but node commonly leaks from $BASE_PATH (/usr/bin), masking
  the forced-missing condition. Force-miss `no-mistakes` instead - a firstmate tool
  never on a standard system PATH - so the diagnostic reliably appears.

bin/fm-lint.sh (ShellCheck 0.11.0) clean.
…gy launches

The captain accepted the no-go: string-scanning a raw launch command can never be
complete against a Turing-complete pane shell (quote concatenation `ag"y"`, brace
`a{gy,}`, alias expansion, generated process substitution, and a wrapper script
that internally execs the binary all defeat a static scan). Move the primary B1
defense to EXEC-TIME interception.

- bin/fm-launch-lib.sh: fm_launch_write_raw_guard writes firstmate-owned refusing
  `cursor-agent`/`cursor`/`agy` shims (exit non-zero) into a guard dir.
- bin/fm-spawn.sh: for a raw launch command, install that guard under
  TASK_TMP/raw-guard and prepend it to the pane PATH before the command runs, so
  ANY spelling that resolves one of those binaries through PATH hits the shim
  instead of the real CLI - uniformly closing quote-concat, brace, alias,
  process-substitution, AND the wrapper-script boundary. Teardown's `rm -rf
  TASK_TMP` cleans it. The sanctioned `--harness` path never routes through the
  guard and reaches the real binary directly.
- The fm_launch_raw_restricted_harness string classifier stays as early,
  spawn-time defense-in-depth (clear pre-launch refusal), not the sole gate.

Documented residual (a PATH shim cannot cover): an absolute-path invocation or a
raw command that first resets PATH bypasses the guard; both are deliberate
circumventions, and closing them needs execve-level interception (LD_PRELOAD/
seccomp) disproportionate to the raw hatch.

Tests: tests/fm-launch-lib.test.sh runs every demonstrated bypass class
(quote-concat agy/cursor, eval+quote-concat, brace, process-substitution, alias,
wrapper-via-PATH) with a fake real binary behind the guard and asserts the real
binary never executes; tests/fm-cursor-agy-adapter.test.sh drives a real raw spawn
and asserts the guard is installed and prepended to the pane PATH. harness-adapters
skill documents the exec-time gate and the residual. shellcheck 0.11.0 clean.
…on A)

Captain chose Option A: MERGE origin/main into the cursor/agy overlay branch
(NOT rebase), preserving the overlay for fast-forward-to-homes landing.

origin/main advanced 4 commits (kunchenguid#895 Calm rendering, kunchenguid#898/kunchenguid#899 operational
markers, kunchenguid#909 canonical operational-input classification). The load-bearing
conflict is kunchenguid#909, which threaded a __OPINPUT__ operational-input encoder and a
__PIBRIEFENV__ Pi env assignment INTO launch_template() in fm-spawn.sh - the exact
function this branch EXTRACTED into bin/fm-launch-lib.sh. Resolution:

- bin/fm-launch-lib.sh (fm_launch_template): threaded kunchenguid#909's
  `"$(__OPINPUT__ encode launch-brief < __BRIEF__)"` into every template (claude,
  codex, opencode, pi, grok, and the new cursor/agy - consistent cross-harness
  operational-input canonicalization) and the `__PIBRIEFENV__` prefix onto the pi
  templates; kept the extraction.
- bin/fm-spawn.sh: kept the extraction (launch_template lives in the lib) and the
  cursor/agy raw-guard/agy-trust/guards; added kunchenguid#909's sq_opinput + PIBRIEFENV
  substitutions using this branch's fm_launch_* helper names.
- tests/fm-calm-pi-extension.test.sh: took origin/main - kunchenguid#895 already re-verified
  Calm on Pi 0.81.1 and pinned it there, superseding this branch's interim
  0.80.10|0.81.1 workaround.
- tests/fm-launch-lib.test.sh: updated the exact-template assertions to the
  kunchenguid#909 __OPINPUT__/__PIBRIEFENV__ form.
- AGENTS.md, harness-adapters SKILL, bin/fm-test-run.sh, fm-pi-watch test:
  auto-merged, both sides preserved (verified).

Merge-affected tests green: fm-operational-input, fm-spawn-dispatch-profile,
fm-pi-watch-extension, fm-cursor-agy-adapter, fm-launch-lib, fm-calm-pi-extension,
fm-transition-lib, fm-agy-trust-lib. shellcheck 0.11.0 clean.
Reconcile the crew-only cursor/agy overlay with seven upstream commits
(a7e01bc..6b0d21d). Merge, never rebase: the overlay must stay
fast-forwardable into detached secondmate homes.

Hand-resolved conflicts:

- bin/fm-spawn.sh launch templates: keep this branch's extraction of the
  templates into bin/fm-launch-lib.sh (fm_launch_template) and its single
  rendering owner (fm_launch_render), rather than main's inline
  launch_template. Template content is at parity with main, including
  kunchenguid#909's __OPINPUT__ operational-input envelope.
- Drop __PIBRIEFENV__ / FM_FIRSTMATE_PI_LAUNCH_BRIEF, which this branch
  still carried from kunchenguid#895. kunchenguid#936 removed that Calm input-reroute binding
  upstream and added a regression test asserting the pi launch command no
  longer exports it, so the removal is adopted here: the placeholder is
  gone from the pi templates and from fm_launch_render's signature.
- tests/fm-backend.test.sh sibling list: take main's ordering; the two
  sides list an identical set.

Also adopt kunchenguid#939's wider ShellCheck source graph, which now lints tests/:
tests/fm-cursor-agy-smoke.test.sh dropped an unused loop variable.

bin/fm-lint.sh is clean and the affected suites pass.
Two suite failures surfaced by merging main, both in contracts main
tightened while this branch was out:

- kunchenguid#939 pins an exact allowlist of tests permitted to carry
  `# shellcheck source=bin/` production context, so the ShellCheck source
  graph stays small. The three new cursor/agy tests had added themselves
  to that set. They do not need production context to lint cleanly, so
  they now stop static source following at /dev/null instead of widening
  the allowlist. bin/fm-lint.sh stays clean.
- The captain-translation contract asserts the launch command carries
  kunchenguid#909's canonical `encode launch-brief` envelope by reading
  bin/fm-spawn.sh. This branch moved the launch templates into
  bin/fm-launch-lib.sh, so the assertion read a file that no longer
  authors them. It now reads both files, covering the launch command
  wherever it is authored.

Not addressed here, and not a regression from this branch:
tests/fm-calm-pi-extension.test.sh fails identically on a clean
origin/main checkout (6b0d21d) with "/calm left the grep row in the
transcript". This branch's .pi/ tree and that test are byte-identical to
main.
feat(harness): add crew-only cursor and agy adapters
… host PATH

The test was at fault, not bin/fm-session-start.sh or bin/fm-bootstrap.sh.

Reproduced on a clean default branch (bash tests/fm-session-start.test.sh):

    not ok - MISSING diagnostic did not appear at all

and, once that site was fixed, the same root cause at a second site:

    not ok - fm-bootstrap.sh's real detect line did not appear verbatim (missing: 'MISSING: node (install:')

Both cases force a bootstrap MISSING diagnostic by deleting the fake `node`
from the fixture bin. That only works when node is absent from the test's
base PATH (/usr/bin:/bin:/usr/sbin:/sbin). On a host that ships node in
/usr/bin, bootstrap's `command -v node` still succeeds, so it correctly
reports nothing missing. Bootstrap's detection is right - node really is
installed - so the fixture, not the script, encodes the stale assumption.

Fixed by giving the fixture a `hide_node` helper that deletes the fake binary
and also masks the lookup through BASH_ENV, the same idiom the Herdr cases in
this file already use to mask tmux. No assertion is weakened or removed: both
MISSING assertions and the diagnostics-before-bulk-context ordering assertion
are unchanged and now actually exercised.

Verified: tests/fm-session-start.test.sh 28/28 pass (exit 0),
tests/fm-bootstrap.test.sh 21/21 pass, bin/fm-lint.sh clean.
No behavior change to session start.
Under a Pi primary, entering away mode left TWO supervision cycles live: the
away-mode sub-supervisor daemon and the Pi watch extension's own arm cycle. The
extension's arm runs `fm-watch-arm.sh --restart`, so it kept taking the watcher
singleton back from the daemon, starving away-mode triage entirely while
injecting ordinary watcher turns into the primary - the behavior recorded live on
2026-07-27.

Reproduced end to end with a real Pi 0.82.1 primary loading the tracked watch
extension, a real away daemon, and the real watcher, against a throwaway
firstmate home on a private tmux socket. Prompts delivered to the primary were
captured by a session-trust extension that aborts before any provider work.

BEFORE - two cycles, daemon starved, primary flooded:
  -- [away entered] state/.afk: present
       1685992 ... bin/fm-supervise-daemon.sh
       1711970 ... bin/fm-watch-arm.sh --restart
       1712043 ... bin/fm-watch.sh        (owned by the extension's arm)
  -- ordinary Pi watcher turns during away routine events: 4
  -- ordinary Pi watcher turns during the away escalation window: 6
  -- away-supervisor escalations delivered to the primary: 0
  daemon log, every cycle:
    watcher non-wake stdout, idling: watcher: already running pid 1685573

AFTER - one cycle while away, daemon triaging, escalation still delivered:
  -- [away entered] state/.afk: present
       1987515 ... bin/fm-supervise-daemon.sh   (its own fm-watch.sh child only)
  -- ordinary Pi watcher turns during away routine events: 0
  -- away-supervisor escalations delivered to the primary: 1
  daemon log:
    self-handle: signal ... -> routine signal: working: another routine step
    escalate: signal ... -> repro-task.status: blocked: needs a captain decision
  after return, exactly one cycle:
       1958195 ... bin/fm-watch-arm.sh --restart
       1958211 ... bin/fm-watch.sh
  -- ordinary Pi watcher turns after return: 3

The extra cycle was armed by the Pi side, not left behind by away mode: the
extension had no notion of the away-mode flag at all.

The extension now observes the durable state/.afk flag. While it is present the
extension arms nothing, retires any arm child it holds, delivers no ordinary
watcher wake, and schedules no retry; a per-generation away poll resumes exactly
one extension-owned cycle once the flag clears. Nothing is lost across the
hand-off because the daemon's own watcher enqueues and triages every event
durably and owns escalation. No process is matched or killed by pattern, and no
non-pi harness path changes - Claude's Stop auto-arm already had the equivalent
gate, and the other adapters are untouched.

Tests:
- tests/fm-pi-watch-extension.test.sh gains a deterministic regression test for
  the dual-supervision condition: arm, raise .afk, assert the arm child is
  retired, assert a further arm reports standby with no second cycle and no
  injected wake, clear .afk, assert exactly one cycle resumes. It fails on the
  pre-fix extension.
- tests/fm-afk-pi-dual-supervision-e2e.test.sh (opt-in, real Pi + real daemon)
  asserts the five acceptance steps and fails pre-fix at "the Pi extension kept
  its own supervision cycle alongside the away supervisor".

Also fixes a pre-existing collation mismatch in bin/fm-test-run.sh and
tests/fm-test-run.test.sh: comm ran without LC_ALL=C over LC_ALL=C-sorted input,
warning "file 2 is not in sorted order" on every coverage-guard run and risking a
wrong comparison under a non-C locale.
…tcher

tests/fm-watcher-lock.test.sh flaked at test_watch_restart_attaches_to_healthy_peer
with:

  not ok - restart did not attach to the verified healthy peer:
           watcher: started pid=N (beacon fresh)

Root cause is a startup race in the fixture, not in the watcher. The case stands
up a TERM-resistant node peer, records it as the lock holder, and expects
fm-watch-arm.sh --restart to attach to it instead of starting a second watcher.
node only becomes TERM-resistant once the interpreter has actually evaluated
process.on("SIGTERM"), and the restart sends TERM within milliseconds of the
launch. When the peer lost that race it died on the default disposition, so the
lock named a dead pid, the fresh child legitimately stole it, and the arm
honestly reported "started" - the fixture failed for a condition it does not
test.

Captured under instrumentation on an isolated reproduction of this case
(3 of 12 iterations):

  iter=8 BAD out='watcher: started pid=1620769 (beacon fresh)|'
     peer=1620694 peer_alive=no
     recorded_identity=linux-starttime=30963098 cmdline-hex=6e6f6465...
     live_identity    =<gone>
     lock_pid=1620769

peer_alive=no with the recorded identity intact rules out the identity/beacon
paths and pins it on the peer dying.

Fix: the peer announces that its SIGTERM handler is installed, and the fixture
waits on that exact condition before recording the lock and arming. No assertion
was loosened, no sleep widened, and no retry was wrapped around an assertion.

Loop evidence - `bash tests/fm-watcher-lock.test.sh`, 50 runs per arm, 6-way
parallel, before-arm run from a pristine `git archive HEAD` copy so both arms are
the same command under the same ambient load:

  before (origin/main):  10/50 FAILED - all 10 the peer-attach race
  after:                  0/50 FAILED

  after, 50 consecutive runs one at a time:  0/50 FAILED

Harness note for anyone re-running this: do NOT launch the loop under `nohup`.
nohup sets SIGHUP to SIG_IGN, an ignored disposition is inherited through fork
and exec, and a shell cannot trap a signal that was already ignored on entry - so
`kill -HUP` inside test_arm_hup_cleans_child_and_temp_output becomes a no-op and
that case times out (124) on every run regardless of the code under test. Early
loops here were launched that way and produced exactly that false signal.

Also verified: bin/fm-lint.sh clean, and all 10 suites selected by
`bin/fm-test-run.sh --changed` pass.
bin/fm-watch.sh ended each supervision cycle with a blind foreground
`sleep "$POLL"`. Bash defers a trapped signal until the running foreground
command returns, so the watcher was deaf to TERM/HUP/INT for the remainder
of the interval, and exit latency tracked FM_POLL exactly. The signal-grace
linger had the same shape.

Measured against an isolated temporary home, signalling a watcher that had
already taken its lock and entered its cycle wait:

  FM_POLL=1   TERM  before 0.54s   after 0.05s
  FM_POLL=5   TERM  before 4.54s   after 0.06s
  FM_POLL=15  TERM  before 14.53s  after 0.06s
  FM_POLL=15  HUP   before 14.52s  after 0.06s

bin/fm-watch-arm.sh --restart allows the outgoing watcher 50 iterations of
`sleep 0.1`, a 5s budget, before forking a replacement. At the 15s default
14.53s exceeded that budget outright, so a restart forked a second watcher
while the first was still alive and still holding the lock; the replacement
then saw a live holder and stood down, and the restart never produced the
fresh watcher it reported wanting.

Run the cycle wait and the signal-grace linger as a tracked background child
and wait on the named pid, so the trap runs the moment the signal lands. The
wait budget, polling cadence, and wake classification are unchanged, and the
EXIT path reaps the sleep child instead of orphaning it for the rest of the
interval (verified: no leaked sleep child after a mid-interval stop).

Liveness and wedge detection are unaffected and covered explicitly: the
heartbeat beacon keeps advancing across cycle waits and goes stale once the
watcher genuinely stops, so a stopped watcher is still detectable.

Known limit, documented at the call site and in the verification record: the
herdr push path waits inside a foreground command substitution and stays up
to FM_POLL deaf. That reader owns a fifo directory and a child reader process
that it removes on its own return path, so interrupting it from the caller
would leak both on every stop. Fixing it needs reader-side teardown.

Regression coverage in tests/fm-watcher-lock.test.sh, all three confirmed
against the unpatched watcher:
- watcher exits on TERM well inside a 10s budget at FM_POLL=60 (fails
  unpatched: "watcher ignored TERM for its whole poll interval")
- --restart hands the lock to a fresh live watcher within the arm's own 5s
  budget, with no overlap between outgoing and incoming (fails unpatched)
- liveness beacon advances while cycling and freezes once stopped (holds
  both before and after; guards against trading exit latency for a watcher
  that keeps its lock while no longer cycling)
…tcher

tests/fm-watcher-lock.test.sh flaked at test_watch_restart_attaches_to_healthy_peer
with:

  not ok - restart did not attach to the verified healthy peer:
           watcher: started pid=N (beacon fresh)

Root cause is a startup race in the fixture, not in the watcher. The case stands
up a TERM-resistant node peer, records it as the lock holder, and expects
fm-watch-arm.sh --restart to attach to it instead of starting a second watcher.
node only becomes TERM-resistant once the interpreter has actually evaluated
process.on("SIGTERM"), and the restart sends TERM within milliseconds of the
launch. When the peer lost that race it died on the default disposition, so the
lock named a dead pid, the fresh child legitimately stole it, and the arm
honestly reported "started" - the fixture failed for a condition it does not
test.

Captured under instrumentation on an isolated reproduction of this case
(3 of 12 iterations):

  iter=8 BAD out='watcher: started pid=1620769 (beacon fresh)|'
     peer=1620694 peer_alive=no
     recorded_identity=linux-starttime=30963098 cmdline-hex=6e6f6465...
     live_identity    =<gone>
     lock_pid=1620769

peer_alive=no with the recorded identity intact rules out the identity/beacon
paths and pins it on the peer dying.

Fix: the peer announces that its SIGTERM handler is installed, and the fixture
waits on that exact condition before recording the lock and arming. No assertion
was loosened, no sleep widened, and no retry was wrapped around an assertion.

Loop evidence - `bash tests/fm-watcher-lock.test.sh`, 50 runs per arm, 6-way
parallel, before-arm run from a pristine `git archive HEAD` copy so both arms are
the same command under the same ambient load:

  before (origin/main):  10/50 FAILED - all 10 the peer-attach race
  after:                  0/50 FAILED

  after, 50 consecutive runs one at a time:  0/50 FAILED

Harness note for anyone re-running this: do NOT launch the loop under `nohup`.
nohup sets SIGHUP to SIG_IGN, an ignored disposition is inherited through fork
and exec, and a shell cannot trap a signal that was already ignored on entry - so
`kill -HUP` inside test_arm_hup_cleans_child_and_temp_output becomes a no-op and
that case times out (124) on every run regardless of the code under test. Early
loops here were launched that way and produced exactly that false signal.

Also verified: bin/fm-lint.sh clean, and all 10 suites selected by
`bin/fm-test-run.sh --changed` pass.

(cherry picked from commit 3db0341)
Both pi tuples in the dispatch config reported no authentication surface, so
rule 2's OpenAI candidate and rule 3's entire MiniMax lane fell through to the
premium default. This was not a credential fault, a quota-axi defect, or a
stale tuple.

The `surface-unresolved` verdict originated in bin/fm-auth-preflight.sh, which
resolved surfaces from a hard-coded harness-to-provider table plus a
`pi:<prefix>` source matcher. It was added in kunchenguid#1349 and deleted four hours
later in kunchenguid#1358, which returned eligibility to open agent judgment resting on
docs/verification/dispatch-auth.md.

That record had no evidence that a provider family absent from quota-axi can
still have a resolved surface, and no record that a harness catalog is itself
surface evidence. With nothing to contradict it, the retired conclusion kept
being re-derived at intake and the zero-premium lane stayed unused.

Pi's installed docs/models.md states that a provider's models stay absent from
--list-models until its auth is configured, so a catalog row proves configured
auth. Verified on Pi 0.82.1 and quota-axi 0.1.16: minimax/MiniMax-M3 is listed
while quota-axi models no minimax provider at all. The surface is resolved by
the catalog; only the quota is unknown, which is disclosed uncertainty.

Records both verified facts, and separates "is the surface resolved" from "is
the quota known" in the decision procedure so an unmodeled vendor can no longer
read as a missing surface. No check is relaxed: this adds a positive evidence
source rather than removing a requirement.

Covered by a case in the live dispatch regression that supplies only the
catalog row and a snapshot without the family. Against the previous procedure
it reports surface=unresolved; with this change it reports surface=resolved.
Sungin Kim and others added 28 commits August 3, 2026 23:22
Reconciles 68 upstream commits with the fork's 17 landed cursor/agy commits.
19 files conflicted; every conflict was resolved by keeping both sides' intent
where compatible and by deferring to the newer owner where one side had moved a
contract.

Notable resolutions:
- bin/fm-dispatch-select.sh and its test: accept upstream's removal of the
  vestigial selector; the fork had only annotated it.
- tests/fm-captain-translation-contract.test.sh and the two source-asserting
  functions in tests/fm-pi-watch-extension.test.sh: accept upstream's removal of
  implementation-source assertions, which the coding guidelines now forbid.
- bin/fm-launch-lib.sh stays the single owner of launch-command construction, and
  upstream's pi-signed and kimi templates plus their model/effort vocabularies are
  ported into it. fm-spawn keeps kimi's spawn-time binary resolution so the
  library's unresolved-placeholder guard still applies.
- bin/fm-test-run.sh: the fork's live-harness-optin exclusion moves into
  list_portable_serial, upstream's single owner of serial membership, so the
  derived CI serial shards stay consistent and the coverage guard proves five
  disjoint partitions.
- bin/fm-teardown.sh: upstream's exact-code error propagation replaces the fork's
  flattened variant, and the agy workspace-trust removal runs after the Herdr
  presentation preflight takes the lock guarding destructive steps.
Adds opt-in GitHub issue traceability. Reconciled against upstream's newer
contract that a ship brief takes its delivery mode as an explicit --mode
argument rather than reading the project registry:

- fm-brief.sh keeps upstream's required, closed-set --mode and gains --issue.
  The branch's registry lookup is dropped; its issue/local-only guard now tests
  the explicit mode.
- --issue argument-shape validation moved ahead of the --mode requirement so a
  malformed issue number reports itself instead of being masked.
- The branch's three tests were ported to name a mode on every ship invocation.
- tests/fixtures/fm-brief-no-issue.sha256 was regenerated against the merged
  tree's own pre-branch output, and the no-issue briefs were verified byte-
  identical before and after this branch, which is what that fixture exists to
  prove.
Both sides had hardened the same MISSING-tool fixture against host PATH leakage:
the fork removes no-mistakes (never present in a system PATH), this branch adds
hide_node. Both statements are kept, because the surrounding assertions require
the no-mistakes MISSING line while the auto-merged run_session_start calls
require the mask this branch produces.
Both sides had reached the same LC_ALL=C comm calls; the branch's fix also
covers the serial-shard comparisons upstream added. Kept the fork's comment
explaining why the collation must match.
The bootstrap-diagnostics trigger list took an addition from each side: the
fork's widened CREW_DISPATCH (invalid or backend mismatch) and this branch's
ENDPOINT_BINDING_MIGRATION. Both are real emitted lines, so both are listed.
Both sides added a window_harness helper to bin/fm-watch.sh, so the merge kept
both and the later one silently shadowed the earlier. That changed what
fm_busy_classify receives for a window with no metadata, from an empty string to
the literal "unknown", with no conflict to review it.

Upstream's definition is retained because its caller was written against the
empty-string form. maybe_native_turnend matches only the exact cursor and agy
tokens, so the crew-only native turn-end path behaves identically either way.

Found by bin/fm-lint.sh (SC2329 on the shadowed definition).
While resolving bin/fm-teardown.sh in the upstream sync, I spliced the per-child
cleanup branch using a context anchor that appears three times in the file. The
edit landed on the first occurrence instead of the intended one and deleted
everything between them, removing six upstream function definitions plus
cleanup_firstmate_home_children:

  cleanup_firstmate_home_children, preflight_firstmate_home_herdr_children,
  teardown_herdr_preflight_target, teardown_herdr_require_prerequisites,
  teardown_herdr_session_lock_held, teardown_release_herdr_locks

Surfaced by tests/fm-backend-autodetect-smoke.test.sh, where teardown died with
"teardown_herdr_preflight_target: command not found".

The file is rebuilt from a fresh three-way merge of the same inputs, with every
conflict resolved as before but with anchors asserted unique first. The result is
verified to define every function either side defines, with no duplicates and
nothing extra; the single exception is registry_home_for_line, which upstream
deleted outright and which has no remaining caller.
tests/fm-backend.test.sh: the legacy fake root that build_old_bin assembles did
not carry fm-agy-trust-lib.sh, which the merged fm-teardown.sh sources for the
cursor/agy workspace-trust cleanup, so the legacy teardown case died sourcing a
file that was never copied. The fork made the same fixture completion in
fm-gotmp.test.sh; this builder was missed because its dependency list grew on the
upstream side. Both fake-root builders now carry the sibling.

tests/fixtures/fm-brief-no-issue.sha256: regenerated against the final tree.
These hashes were pinned when fm/fm-issue-lifecycle-wiring merged, and
fm/fm-brief-status-protocol-gaps subsequently changed the crewmate scaffold's
status protocol, which legitimately changes ship and scout brief bytes. Only
those six entries moved; the two secondmate charters are unchanged, which matches
a crewmate-scaffold-only change. The property this fixture exists to protect -
that a no-issue brief is unaffected by the issue feature - was verified byte-for-
byte when that branch merged and is still asserted by
test_issue_traceability_is_strictly_opt_in.
The fork's cursor/agy test predates three upstream changes and could no longer
reach the refusals it asserts:

- Ship spawns now require explicit --mode and --yolo, and both checks run before
  the crew-only and herdr-only gates. The ship cases now name both; the
  --secondmate rows in the raw-bypass table deliberately pass neither, because a
  secondmate records its own fixed posture and refuses those flags.
- fm/fm-endpoint-binding-migration refuses an endpoint record with no exact task
  binding, so the forced-child fixture now records endpoint_task_id.
- That same fixture recorded its child on herdr, which now additionally has to
  clear teardown_herdr_preflight_target's structured pane inspection. The case
  is about cleanup_firstmate_home_children refusing when a child's firstmate-owned
  agy trust cannot be removed, which runs before any backend-specific work and is
  backend-independent, so the child is recorded on tmux instead. Faking herdr's
  pane protocol here would prove nothing about trust cleanup; real herdr endpoint
  behaviour stays covered by the real-herdr-gated lane.

The gates themselves are unchanged; only the fixtures were conformed.
…ncile

Sync the fork with upstream and land the ten reviewed changes
* feat(bin): preflight remote runtime tool paths (kunchenguid#1623)

* feat(bin): widen the remote runtime PATH and add a remote doctor preflight

The fixed remote entrypoint hard-coded a four-directory PATH, so a remote
account whose tools live under nix or a per-user profile could not run basic
Firstmate work without a login shell. The entrypoint now composes its child
PATH from the code root's bin, the account's ~/.local/bin, the common
package-manager directories that actually exist on the host, and the portable
system tail, deduplicated and in a fixed order, still under env -i with the
same variable allowlist and no shell command string.

fm-remote-doctor.sh reports that exact PATH by inheriting it from its own
entrypoint launch rather than recomposing it, so the ordering keeps one owner.
It is read-only, reports where each required and optional tool resolved, and
exits non-zero naming every required tool that did not. Remote seeding runs it
as a preflight before anything is created on the host and restores the registry
when it fails.

* no-mistakes(review): Harden remote git authorization and missing-tool diagnostics

* no-mistakes(document): Document remote PATH doctor and safe shims

* no-mistakes(lint): Fix ShellCheck findings in remote path tests

* no-mistakes(lint): Suppress exported fixture's false-positive ShellCheck warning

* feat(bin): detect documentation-vault drift across project clones

A project's knowledge vault could fall dozens of commits behind with nothing
surfacing it, because noticing depended entirely on someone remembering to look.
fm-vault-drift.sh closes that gap as a cheap, read-only inspection of every
registered project clone, relayed by bootstrap as a VAULT_DRIFT diagnostic line
in both normal and detect-only sessions.

It is detection only: a vault is curated knowledge, so an automated writer would
manufacture exactly the stale-but-confident prose the check exists to catch.

The two vault shapes are told apart by inspecting the clone, never by trusting
registry prose, because they need different remedies: an in-repo vault a
crewmate can refresh in its own worktree, and an external symlinked vault living
in a separate repo that an isolated project worktree structurally cannot write.
An absent or broken link reports distinctly from staleness, since "drift cannot
be measured at all" is the failure that hides staleness rather than a form of it.
Staleness carries the commit count behind and the drift window, both derived
from commit timestamps so a run is deterministic.

A directory that merely shares the name vault/ without the OKF bundle marker is
never reported, so test fixtures and sample trees raise no false alarm.

Also fixes two pre-existing test problems found while validating: the
session-start ordering and composition cases forced a MISSING diagnostic by
removing node, which a host-installed /usr/bin/node silently satisfied again,
and the test-runner coverage guard compared LC_ALL=C-sorted files with an
ambient-locale comm, so it failed outright in any non-C locale.

* no-mistakes(review): Fix post-sync vault drift and marker validation

* no-mistakes(review): Reject project repository as external vault target

* no-mistakes(document): Document vault drift detection and configuration

* fix(tests): keep vault-drift fixtures deterministic under Git 2.54

CI runs on Git 2.54, which enables automatic maintenance by default. Against
this file's rapid fixture commits that maintenance races the index and deletes
loose objects the pending commit still references, so the fixture dies with
"invalid object ... Error building trees" partway through the 52-commit case
and the stale external vault is never measured.

Reproduced at 8/8 under Git 2.54 and 0/8 after disabling maintenance on the
fixture repos; Git 2.51 and 2.43 never trip it, which is why it passed locally
and only surfaced in CI. gc.auto=0 does not cover it - this is the maintenance
path, not the gc one.

The detector is unchanged: this only makes fixture history deterministic.

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Sungin Kim <sunginapp@gmail.com>
* fix(afk): deliver away-mode escalations to an idle primary

The away-mode sub-supervisor buffered escalations correctly and replayed
every one on return, but could not inject into the primary session for up
to eleven hours at a time - three overnight wedges of 7.8h, 11.0h, and
10.1h. Durability held; responsiveness did not.

The recorded hypothesis was that a long mid-turn Claude pane is simply
outside what the injection guard was designed for. The daemon log
disproves it: all three wedges are dominated by `state=pending`
deferrals (2792 on the incident day alone), while the busy branch that a
genuine mid-turn pane takes fired 11 times in the entire log history.

Root cause is a misread of an idle, injectable composer. Claude 2.x
renders its EMPTY composer row as `❯` + U+00A0 NO-BREAK SPACE, and bash's
`[:space:]` trims treat no non-ASCII blank as whitespace under any locale
the fleet runs, so the pad survived every trim as apparent typed content
and classified `pending`. The injector types only into an affirmatively
`empty` composer, so it deferred forever against a pane that was ready
the whole time. Reproduced live on the primary pane, which read
busy=idle composer=pending before this change and composer=empty after.

fm_composer_blank_normalize folds non-ASCII blank and zero-width
characters to an ASCII space before the verdict, in the shared owner so
every adapter gets it. The fold is safe in one direction only: every
character it folds is invisible, so any visible byte still reads
`pending` and the dead-shell refusal is untouched.

The leading prompt glyph is now removed as a literal prefix. Under an
exported C/POSIX locale `${content#?}` drops one BYTE, leaving the two
trailing bytes of a multibyte glyph behind as spurious content - a second,
independent deferring condition, verified to turn an idle placeholder
composer from `empty` into `pending`.

The injection guard itself is unchanged: it was fed a wrong verdict, not
reasoning wrongly. A genuinely mid-turn primary still defers on the busy
guard, bounded by one turn and self-clearing, with the max-defer wedge
alarm still covering a pathological one.

Buffering and replay-on-return are untouched, and nothing is delivered
that was not delivered before.

* no-mistakes(review): Document mid-turn escalation delivery limitation

* no-mistakes(document): Document composer padding and mid-turn delivery limits

---------

Co-authored-by: Sungin Kim <sunginapp@gmail.com>
* fix(watch): suppress parked-decision stale wakes on pane repaint

A crew that correctly parks on a keyed `needs-decision:` is genuinely not
working, so `crew_is_provably_working` rightly refuses to absorb its stale
pane. The one-shot stale suppressor was keyed on the pane content hash, but
an idle harness pane is not static - its footer carries a live context
percentage and quota readout - so every repaint changed the hash and
re-surfaced the identical, already-escalated waiting state as a fresh stale
wake. One parked pane escalated twenty-two consecutive times while healthy,
each costing a full handling turn.

Key the terminal stale path's suppressor on the open-decision identity from
`status_open_decisions` (bin/fm-classify-lib.sh's existing keyed fold)
whenever the captain-relevant status is a still-open keyed decision, so a
repaint alone is silent. Detection is not weakened: a new status line, a new
decision key, or the decision resolving all change the open set and surface
normally, and a parked crew whose backend confidently reports its agent dead
still escalates through the shared wedge timer. Only a confident `dead`
verdict counts, so an ambiguous or unreadable backend read cannot manufacture
a wedge alarm.

Tests cover all three properties and fail against the previous behavior.

* no-mistakes(review): Captain, reconcile surfaced open-decision markers

* no-mistakes(document): Clarify parked-decision stale wake documentation

---------

Co-authored-by: Sungin Kim <sunginapp@gmail.com>
* fix(crew-state): keep the authoritative source through a fix round

fm-crew-state.sh reported `unknown / none / no current-state source
available` for crews that were demonstrably mid-validation. Because
fm-classify-lib.sh's crew_is_provably_working() reads the same line, a
working crew stopped being provably working, so fm-watch.sh could not take
its absorbed-stale path and every idle repaint raised a fresh possible-wedge
escalation - precisely during the longest phase of a run.

Three distinct mechanisms produced that one symptom. All three are fixed.

1. A lookup FAILURE read as a run ABSENCE. nm_run swallowed the exit status
   with `|| true`, so a `no-mistakes` call killed by its own timeout under a
   saturated daemon was indistinguishable from "this branch has no run".
   The status is now propagated (124 on timeout, 127 when the call cannot be
   bounded at all), and a failure degrades to the crew's last observed
   run-step under a new source, `run-step-degraded`, recorded in
   state/<id>.run-step. Empty stdout counts as failure: verified against the
   installed CLI that `axi status` answers a branch with no run of its own
   with some other branch's run, so it has no empty "nothing to report".

2. The code-identity skip fired during a fix round. When a review finding is
   answered `--action fix` the pipeline commits the fix, the tip advances past
   the sha the run recorded, and the run that was authoring those very commits
   stopped matching its own worktree. An actively-executing run now keeps
   attribution across that advance; a parked run commits nothing, so a tip that
   moved past it is still local work outside the run and still invalidates.

3. The run's head was not an object in the worktree at all. Verified live on
   arkrh-corpus-types-unbound-unconstructible: `axi status` in the crew's own
   worktree returned that crew's own branch, status running, parked at
   fix_review with 3 findings and branch_sync.state pipeline_owned, but its
   head 1c17405c failed `git rev-parse --verify` ("malformed object name")
   because the pipeline commits in a copy the worktree never fetched. The
   lookup completed in under a second and the tip had not moved, so this is
   neither mechanism above. nm_head_relation now separates `unresolved` from
   `diverged`, and a LIVE run answered for this branch may be attributed
   despite an unresolvable head.

Against that live task the shipped reader could not attribute at all and fell
through to the pane; this branch reports `working / run-step / validating
(running)`. Read-only check against a temp state dir, never the live home.

Not weakening wedge detection is the constraint that shapes all three, and
every widening is paired with a guard test that must keep passing:

- Nothing is degraded for a crew never seen validating: no record, no replay.
- The degrade is age-bounded (FM_CREW_STATE_DEGRADED_MAX_AGE, default 900s,
  0 disables), so a permanently unreachable daemon stops absorbing suspicion.
- A gone endpoint and an exact busy verdict both outrank a replayed record;
  live evidence always beats memory.
- A completed lookup that finds no run never degrades - absence stays absence.
- Terminal runs are refused in every widened case, and the historical runs
  listing never earns the pipeline-owned benefit that only a branch-scoped
  `axi status` answer does.
- run-step-degraded is deliberately excluded from fm-fleet-snapshot.sh's
  decision-clearing sources: good enough for wedge triage, never good enough
  to clear a captain decision.

Tests: 5 reproductions fail on the current code and pass here; 7 safety tests
pass on both sides, which is what makes them guards. Verified per-test in both
directions rather than only as a suite.

Pre-existing and out of scope, each already queued separately and confirmed
failing on the base branch with these changes stashed:
tests/fm-calm-pi-extension.test.sh, tests/fm-session-start.test.sh,
tests/fm-test-run.test.sh. The rest of the non-e2e suite passes (76 files).

* no-mistakes(review): Invalidate stale run-step records after confirmed absence

* no-mistakes(review): Invalidate cached run-step on same-branch rejection

* no-mistakes(review): Preserve cached run-step for inconclusive missing heads

* no-mistakes(review): Persist CI-ready verdicts before terminal emits

* no-mistakes(review): Reject coarse run-behind rows without authoring proof

* no-mistakes(document): Document degraded crew-state attribution

---------

Co-authored-by: Sungin Kim <sunginapp@gmail.com>
* docs: record GPU reranker endpoint

* no-mistakes(review): Captain, redact host-specific reranker verification evidence

* no-mistakes(document): Remove task-specific reranker documentation references

---------

Co-authored-by: Sungin Kim <sunginapp@gmail.com>
* feat(gbrain): add local embedding endpoint artifacts

* no-mistakes(review): Fix GBrain endpoint persistence and configuration, captain

* no-mistakes(review): Captain, document GBrain endpoint environment requirements

* no-mistakes(document): Clarify GBrain verification scope

---------

Co-authored-by: Sungin Kim <sunginapp@gmail.com>
* Resolve work items against each project's own issue tracker

Firstmate manages projects across several forges and hosts, but a task's
issue identity was a bare number with no project attached. The merge path
then closed that number against the owner/repository parsed out of the PR
URL, so any project whose code and issues live in different places had its
bookkeeping addressed to the wrong tracker, silently.

Give the project registry a tracker declaration and resolve every reference
through it:

- data/projects.md gains a tracker=<forge>:<host>/<path> token inside the
  existing bracket annotation, with tracker=none as an explicit "no tracker"
  distinct from an absent declaration. The token never disturbs delivery
  posture parsing, and the tracker is never inferred from a git remote, a
  clone directory name, or a PR URL.
- bin/fm-issue-lib.sh owns the declaration and the accepted reference forms:
  a full URL, a <forge>:<url> prefixed URL for the self-hosted shape several
  forges share, <owner>/<repo>#<n>, and a bare #<n>. A form that needs a
  declaration and has none is refused with an actionable reason.
- bin/fm-issue-ref.sh resolves references at intake, the same way delivery
  mode and yolo are resolved once and passed on explicitly. A task may carry
  several references or none; one unresolvable reference refuses the set.
- fm-brief.sh --work-item takes only a resolved reference and never reads the
  registry. fm-spawn.sh records work_item= lines in task metadata and upgrades
  a legacy bare issue marker through the declared tracker, reporting rather
  than guessing when a project declares none.
- fm-pr-merge.sh closes a recorded GitHub work item in the repository that
  record names. Only the legacy bare number still falls back to the PR's
  repository, which is all a bare number can mean.
- bin/fm-issue-status.sh adds optional title and open/closed enrichment for
  GitHub and Gitea, with GitLab reporting that it has no adapter. Every
  failure degrades to the link plus a reason and still exits 0, results are
  cached, and live lookups are spaced per host.

Per-host credentials live in config/forge-tokens/<host> at mode 0600, refused
if stored more loosely, absent from the inherited-config allowlist, and passed
to curl through a stdin config so they never reach process arguments.

Cross-forge fixtures cover a GitHub project, a Gitea project, a renamed
repository whose clone directory disagrees with its tracker, a mirrored
project whose git remote points elsewhere, an undeclared tracker, tracker=none,
a malformed declaration, malformed references, and an unauthorized host.

* no-mistakes(review): Captain, harden work-item linkage and status enrichment

* no-mistakes(review): Harden host-safe issue enrichment and write-back

* no-mistakes(review): Serialize status rate limiting and refuse ambiguous trackers

* no-mistakes(review): Release rate-limiter claims only under proven ownership

* no-mistakes(review): Drop claim protocol for best-effort cached rate limiting

* no-mistakes(document): Document project-scoped tracker linkage

---------

Co-authored-by: Sungin Kim <sunginapp@gmail.com>
* Record a durable outcome manifest and extend the fleet snapshot contract

A finished task used to leave nothing structured behind. Teardown removes
state/<id>.meta, and the backlog's Done section is pruned to a recent window,
so once a task was cleaned up there was no canonical record of what it was,
what it shipped, or which session ran it. Nothing downstream could attribute
its usage, ingest its outcome, or place it in history.

Introduce three data contracts and project all three through the read-only
fleet snapshot:

- data/<id>/outcome.json (fm-outcome-manifest.v1) is the canonical completion
  manifest. Teardown publishes it atomically before removing the volatile
  records it is composed from, and refuses the cleanup when publication fails
  rather than erasing a task it could not archive.
- data/<id>/work-items.json (fm-work-items.v1) carries forge- and host-agnostic
  work-item references, several per task or none, each marked as declared at
  intake or derived from a PR, with nullable enrichment consumers must render
  without. This owns storage and transport only; population and per-forge
  resolution stay with the project-issue-linkage work.
- state/<id>.pr-status (fm-pr-status.v1) holds one normalized review, check,
  and mergeability observation per PR, refreshed only by fm-pr-status.sh so
  read-only consumers never call a forge.

fm-fleet-snapshot.v1 grows additively, so existing v1 consumers keep working
unchanged: task model and effort, the last task-event timestamp and its age,
the parsed PR identity with its normalized state, work-item references, a
computed card column resolved against a published and tested precedence
ladder, watcher heartbeat age with away-mode state, and manifest-backed
durable history that outlives teardown.

Secret safety is enforced rather than described. A fixed recursive key
allowlist gates manifest publication, so a field can only be emitted by being
declared in one place, and every free-text value is stripped of control
characters, collapsed to one line, and length-capped.

jq becomes a universal bootstrap requirement: the snapshot, the dispatch
profile reader, and now the manifest all depend on it, and without it a home
cannot archive a finished task.

* Publish the manifest before a retired secondmate's home is removed

A retired secondmate's state and data directories can live inside the home
teardown removes, which is exactly the shape the remote parent-route shadow
uses. Publishing the manifest after that removal left the writer with no task
metadata to compose from, so the remote lifecycle end-to-end run failed on a
retirement that should have succeeded.

Move the publication to the last point where every source record still exists:
after the final refusal gate and the confirmed endpoint close, before the home
removal. Pin it with an assertion that a retired secondmate leaves a manifest
recording its retirement in the parent's durable history.

Also give the gate-refusal teardown fixture its own data directory so the suite
stops writing a manifest into the repository's own gitignored data/.

* Give the backend conformance fixture the manifest writer it now needs

The old-versus-new teardown conformance case rebuilds a bin/ tree from a fixed
sibling list, and teardown now shells out to fm-outcome-manifest.sh. Without
the writer and its library in that tree the rebuilt teardown could not publish
a manifest and refused a cleanup that should have succeeded.

* no-mistakes(review): Harden outcome contracts and preserve cache identity

* no-mistakes(review): Normalize GitLab approvals and bound outcome validation

* no-mistakes(review): Enforce outcome contracts and approval precedence

* no-mistakes(document): Correct fleet data contract documentation

* no-mistakes(document): Align fleet documentation wording

* Stop the manifest validator from refusing tasks it should archive

CI caught two ways the hardening round's value validation was stricter than
the rest of firstmate, and because teardown refuses when a manifest cannot be
published, each one made a real task impossible to clean up.

The task-id rule invented a 64-character cap and a bounded character pattern.
bin/fm-pr-lib.sh's fm_task_id_path_safe has no length cap, so a legacy id the
whole rest of the system accepts could not be archived and its teardown
refused. The manifest now mirrors that predicate exactly: the archive records
the ids that exist rather than policing them.

The mode field was validated against a closed delivery-mode vocabulary, but
mode is provenance copied verbatim from task metadata, and task kinds appear
there in practice alongside delivery modes. A scout task therefore could not be
torn down at all. mode now gets the same bounded type and length treatment as
project, harness, and model; enums stay only on the values this contract
generates itself.

Also bump the snapshot test count CI pins for the stock-macOS-Bash job from 15
to 19, which this branch's four new fleet-snapshot cases changed. The Bash 3.2
parse sweep in that job passed unchanged.

---------

Co-authored-by: Sungin Kim <sunginapp@gmail.com>
* feat(bin): verify the model a dispatched worker actually ran on

`model=` in state/<id>.meta is what firstmate REQUESTED at intake, sometimes
after a deliberate quota-balanced choice. Nothing verified what RAN. A worker
served below its dispatched tier still reports done and the record still reads
as the intended model, so later quota-aware dispatch built on that record is
fiction.

Placement, and why not the alternatives:

- Not at spawn time: the worker has produced no turn yet, so there is nothing
  to compare. Verification is necessarily after the fact.
- Not a PostToolUse guard: that surface exposes `resolvedModel` for the
  harness's OWN delegation tool, which a firstmate primary already denies
  (docs/subagent-guard.md), and it never observes a bin/fm-spawn.sh dispatch -
  the record actually at risk.
- Not folded into bin/fm-crew-state.sh: that helper owns one contract,
  reconciling current run state. Model provenance is orthogonal to it.

Shipped as bin/fm-model-verify.sh, a read-only verifier surfaced through the
fleet snapshot firstmate already reviews every heartbeat. Evidence is the
harness's own session transcript: the runtime writes it, the agent never
authors it, and it records the model that served each assistant turn. Asking
the worker instead would query the one party that cannot see the answer.

It fails loudly rather than reporting compliance. `match` requires that a model
was actually read and compared. Every path that cannot read the truth - no
evidence adapter for the harness, evidence unlocatable or unreadable, jq
absent, evidence unattributable to this dispatch - ends in `unverifiable`,
never in silence. `pending` and `unpinned` are explicit no-verdict outcomes and
never render as verified.

bin/fm-spawn.sh now records `spawned_at=`. A worktree from a reusable pool can
still carry a previous occupant's transcripts; without that anchor the two
occupants cannot be told apart, which was observed live across the development
fleet.

The view raises a Model Routing section only when a worker did not provably run
as dispatched, so correctly routed work renders exactly as before.

The composition the source investigation flagged as unverified - that a
PostToolUse hook on a delegation call sees `resolvedModel` - was confirmed
empirically before any of this was built. It holds. That evidence, the
transcript ground truth, the path encoding, and the live observation are
recorded in docs/verification/model-verification.md.

Also fixes an unrelated pre-existing break in the test runner's coverage guard:
its lists are built with `LC_ALL=C sort`, but its `comm` calls ran under the
ambient locale, which reads a C-sorted list as unsorted, exits nonzero, and
lets `set -eu` abort the guard with nothing but comm's warning.

* no-mistakes(review): Harden model verification across spawn and teardown

* no-mistakes(review): Harden model verification across dispatch and teardown

* no-mistakes(review): Pin Claude launches to canonical evidence stores

* no-mistakes(review): Reject newline-bearing model evidence stores

* no-mistakes(document): Document teardown model-verification gate

* no-mistakes(lint): Fix ShellCheck empty suffix declaration

* fix(bin): scope the terminal model-routing refusal to verifiable dispatches

Full CI caught what the pipeline's focused test step did not: the terminal
model-routing refusal blocked legitimate teardown across five lanes - scouts
with a report present, empty secondmate homes, zellij ghosts, Orca scouts. A
harness with no evidence adapter can never produce a verdict, so refusing
whenever one was absent made non-forced cleanup permanently impossible for
every non-claude worker. That is a fleet-wide regression, not the boundary the
refusal was meant to draw.

The verdict is now ALWAYS surfaced before cleanup, so no worker's model
provenance is discarded unseen. Only the refusal is conditional: a mismatch
always blocks, and an absent verdict blocks only when the dispatch was
verifiable in principle - a harness with an evidence adapter and a pinned
model. `--force` keeps its existing discard authority, and only non-forced
teardown refuses.

Also fixed, all found by full CI:

- The macOS lane is a pinned test-count guard, not a portability check: every
  test passes under stock bash 3.2. Four tests were added to the
  snapshot/fleet-view suite, so the pinned count moves 15 -> 19 rather than the
  suite being trimmed to fit a stale guard.
- Three watermark cases used the bare spawn helper, so ship spawns correctly
  refused for a missing --mode. They now use the ship helper like their peers.
- fm-teardown.sh gained a dependency on fm-model-verify.sh, which two fixtures
  that assemble a minimal bin/ did not provide. Both now supply it, so a
  genuinely missing verifier still refuses loudly rather than being papered
  over in the fixture.
- The Herdr projection comparison masked only container ids, so per-dispatch
  identity fields that can never agree between two spawns read as a projection
  difference. Those fields are masked too, keeping the comparison on what it
  exists to test.

Terminal-mode boundary is now pinned in both directions: no block for a harness
that can never produce a verdict or a dispatch that pinned no model, always a
block on a mismatch.

* test(ci): pin the snapshot guard to the merged suite size

Rebasing onto main composed two additive test sets in the snapshot/fleet-view
suite - the landed fleet-telemetry work and model verification - so the stock
macOS lane's pinned count moves 19 -> 23 to match what the suite now runs.

---------

Co-authored-by: Sungin Kim <sunginapp@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant